Papers with diagnostic tests

6 papers
Improving Precancerous Case Characterization via Transformer-based Ensemble Learning (2022.emnlp-industry)

Copied to clipboard

Challenge: Application of natural language processing (NLP) to cancer pathology reports has been focused on detecting cancer cases, ignoring precancerous cases.
Approach: They developed transformer-based deep neural network NLP models to perform the CRC phenotyping with the goal of extracting precancerous lesion attributes and distinguishing cancer and precancirous cases.
Outcome: The proposed model achieves 0.914 macro-F1 scores for classifying patients into negative, non-advanced adenoma, advanced adénoma and CRC.
SODAPOP: Open-Ended Discovery of Social Biases in Social Commonsense Reasoning Models (2023.eacl-main)

Copied to clipboard

Challenge: Existing diagnostic tests for detecting social biases in NLP models only detect stereotypic associations pre-specified by the designer.
Approach: They propose an approach for automatic social bias discovery in social commonsense question-answering by substituting names associated with different demographic groups and generating many distractor answers from a masked language model.
Outcome: The proposed approach uncovers model’s stereotypic associations between demographic groups and an open set of words.
Finding Culture-Sensitive Neurons in Vision-Language Models (2026.eacl-long)

Copied to clipboard

Challenge: Vision-language models struggle on culturally situated inputs, study shows . despite impressive performance, many VLMs struggle on such culturally grounded inputs .
Approach: They propose a new margin-based selector to identify neurons associated with cultural selectivity . they also introduce a model-dependent decoder to identify such neurons .
Outcome: The proposed model outperforms probability- and entropy-based methods in identifying neurons associated with cultural selectivity.
How Reliable are Model Diagnostics? (2021.findings-acl)

Copied to clipboard

Challenge: Contemporary statistical models trade off interpretability and simplicity for powerful parameterizations and inductive biases, enabling impressive performance.
Approach: They examine three recent models and find they are not yet reliable . they also formulate recommendations for practitioners and researchers .
Outcome: The proposed models are not as reliable as previously assumed, the authors argue . their findings suggest that they are needed for improving models and training setups .
Probing for Semantic Classes: Diagnosing the Meaning Content of Word Embeddings (P19-1)

Copied to clipboard

Challenge: Empirical analysis of word embeddings of ambiguous words is limited by the small size of manually annotated resources and by the fact that word senses are treated as unrelated individual concepts.
Approach: They present a large dataset based on manual Wikipedia annotations and word senses, where word sense from different words are related by semantic classes.
Outcome: The proposed method can predict whether a word is single-sense or multi-sensor, if the sense is frequent, and it can predict rare senses.
Investigating the Effect of Pre-finetuning BERT Models on NLI Involving Presuppositions (2023.findings-emnlp)

Copied to clipboard

Challenge: a study of presupposition, discourse and sarcasm suggests that pre-finetuning can improve models' performance on presimplified cases.
Approach: They propose to leverage the connection between presupposition, discourse and sarcasm to improve models' performance.
Outcome: The proposed model improves on cases involving presupposition by pre-finetuning on additional tasks and datasets.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations